[FEAT]: Add durable terminal evaluation contract - #146
Spencer Schoenberg (spencrr) merged 5 commits into
Conversation
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
3cc7cc8 to
fb90c72
Compare
c3b5f80 to
9b7c6a6
Compare
9b7c6a6 to
ef98095
Compare
There was a problem hiding this comment.
🟡 Changes recommended
A huge terminal confidence value can raise an uncaught overflow and abort controller-side xdist merging.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds durable terminal-evaluation and population provenance across core results, reporting, and xdist transport.
Changes:
- Adds terminal evaluation, trace-end reason, evaluation-purpose, and resolver APIs.
- Centralizes population validation.
- Extends persistence, truncation handling, tests, and documentation.
File summaries
| File | Description |
|---|---|
rampart/core/_population.py |
Adds shared population validation. |
rampart/core/types.py |
Adds provenance enums and turn purpose. |
rampart/core/result.py |
Extends results and verdict resolution. |
rampart/core/execution.py |
Validates trial parameters early. |
rampart/core/__init__.py |
Exports new core APIs. |
rampart/pytest_plugin/_xdist.py |
Transports and validates provenance. |
rampart/reporting/json_file.py |
Serializes terminal metadata. |
rampart/probes/_single_turn.py |
Clarifies turn-limit behavior. |
rampart/probes/_factory.py |
Updates probe API documentation. |
rampart/attacks/_xpia.py |
Clarifies XPIA turn limits. |
rampart/attacks/_factory.py |
Updates XPIA API documentation. |
tests/unit/core/test_types.py |
Tests enums and turn validation. |
tests/unit/core/test_result.py |
Tests result contracts and resolvers. |
tests/unit/core/test_execution.py |
Tests early trial validation. |
tests/unit/pytest_plugin/test_xdist.py |
Tests transport and truncation. |
tests/unit/pytest_plugin/test_xdist_aggregation.py |
Tests worker/controller persistence. |
tests/unit/reporting/test_json_file.py |
Tests JSON provenance output. |
docs/api/core-types.md |
Documents new core APIs. |
docs/usage/results-and-reporting.md |
Explains result provenance. |
docs/usage/xdist.md |
Documents envelope behavior. |
docs/probes/behavioral.md |
Updates probe turn-limit semantics. |
docs/attacks/xpia.md |
Updates XPIA turn-limit semantics. |
Review details
- Files reviewed: 22/22 changed files
- Comments generated: 1
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
ef98095 to
acbf019
Compare
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
There was a problem hiding this comment.
🟡 Changes recommended
Extreme malformed integers can escape the xdist fail-closed error boundary.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
Suppressed comments (1)
rampart/pytest_plugin/_xdist.py:998
- A sufficiently large integer overflows
float(), then its!rconversion can itself raiseValueErrorunder Python's integer-string digit limit. That exception escapes instead of becomingWorkerOutputError, so malformed worker data can abort the pytest hook rather than marking the run incomplete. Render the value through the existing safe string helper.
msg = f"Confidence could not be converted to float: {raw_confidence!r}."
- Files reviewed: 23/23 changed files
- Comments generated: 3
- Review effort level: Balanced
Keep numeric overflow and evaluation text conversion failures inside the worker-output error boundary so controllers mark runs incomplete and preserve earlier results.
Require explicit response scopes and distinguish online evidence from terminal verdict input without an alias. Apply canonical policies consistently to shared evaluation schemas and version the tightened population contract. Migrate callers and document direct pre-1.0 API replacement while preserving fail-closed records.
acbf019 to
226cf4e
Compare
Use final_trace_evaluation consistently in Result, canonical records, JSON reports, and xdist transport. Update examples and round-trip tests without retaining an alias for the earlier spelling.
Nina Chikanov (nina-msft)
left a comment
There was a problem hiding this comment.
Thanks for addressing my previous comments! This round of review focused on the major bump to the schema and associated changes. I don't think the deprecation change needs a lot of discussion but is worth aligning on with Behnam at the very least (since he helped push through #185) in a quick call or in PL.
Validate inline migration instructions instead of document paths. Keep migration mechanics generic, align schema support with the project-wide deprecation policy, and focus xdist documentation on user-visible guarantees.
Nina Chikanov (nina-msft)
left a comment
There was a problem hiding this comment.
Thanks for the changes!
ecd2bb7
into
microsoft:main
Description
Adds the per-execution provenance needed before final-trace verdict cadence changes.
Result.terminal_evaluationstores the evaluator output for the terminal trace,Result.trace_end_reasonrecords why trace production ended, andTurn.eval_purposeidentifies online stop checks.Result.turn_evaluationsmakes the online evidence boundary explicit whileeval_resultsremains a compatibility view.The direct attack and probe resolvers require one evaluation and reject unknown runtime outcomes instead of falling through. Population provenance now shares validation across
PopulationRef,PopulationResult, andexecute_trials_async, with invalid thresholds rejected before an execution factory runs.This PR also includes the persistence work previously split into #147 so the contract cannot land without transport support. The current xdist v2 envelope and JSON report carry terminal evaluation, trace-end reason, turn purpose, and population provenance together. Malformed worker data fails closed, including overflowing confidence values, and oversized results produce bounded incomplete markers while retaining population provenance when it fits.
Breaking changes
None for valid callers. Invalid population provenance and malformed evaluator outcomes now fail early instead of being accepted or falling through.
Checklist
pre-commit run --all-filespassesValidation: 1,122 unit tests pass. Strict documentation build and all pre-commit checks pass.